Back

Bioinformatics Advances

Oxford University Press (OUP)

Preprints posted in the last 90 days, ranked by how well they match Bioinformatics Advances's content profile, based on 203 papers previously published here. The average preprint has a 0.19% match score for this journal, so anything above that is already an above-average fit.

1
EcoXAI: Autonomous Agentic Ecosystem for Explainable Artificial Intelligence and Biomedical Discovery

Matsumoto, N.; Choi, H.; Freda, P. J.; Hernandez, M. E.; Wang, Z. P.; Moore, J. H.

2026-07-13 bioinformatics 10.64898/2026.07.08.737358 medRxiv
Top 0.1%
30.8%
Show abstract

MotivationAs biomedical datasets and knowledge graphs continue to grow in size, complexity, and heterogeneity, navigating and extracting actionable insights from them presents a major bottleneck for researchers. There is a clear need for autonomous analytical solutions that can utilize recent advancements in agentic AI such as agent harnessing and loop engineering without introducing hallucination or workflow fragmentation. Researchers, regardless of technical expertise, need tools that streamline complex data analysis and deliver meaningful, actionable insights grounded in both data and established biomedical knowledge. EcoXAI addresses this by introducing a modular, customizable, containerized multi-agent system that structures analysis into explicit pipeline execution stages, lowering the computational barrier for clinical and translational researchers. ResultEcoXAI replaces monolithic AI text interfaces with an autonomous execution-driven framework with specialized bioinformatics agents for delivering proactive, data-driven insights grounded in established biological knowledge. Unlike purely LLM-driven or less integrated AI solutions prone to hallucinations or biologically implausible outcomes, EcoXAIs multi-agent framework, which leverages modern agentic management and explicit knowledge graph integration, provides greater transparency and verifiability in its reasoning. In our use case in drug repurposing for Alzheimers Disease, EcoXAI evaluated 103 drug candidates and identified 79 novel candidates whose predictive models exceeded a randomized baseline, including the CCR5 antagonist Maraviroc, whose generated hypothesis was subsequently supported by the literature. These results demonstrate the potential of knowledge graph-grounded AI agents to accelerate hypothesis-driven biomedical research. Availability and implementationEcoXAI is available on GitHub at: https://github.com/EpistasisLab/EcoXAI. Contactjason.moore@csmc.edu

2
Deep-Interact Studio: An Interactive Deep Learning Model Building Platform for Biomolecular Interaction Prediction

Sarkar, D.; Bardhan, K.; Sarkar, C.

2026-07-07 bioinformatics 10.64898/2026.07.02.736034 medRxiv
Top 0.1%
22.9%
Show abstract

Motivation: Deep learning has rapidly become essential for predicting biomolecular interactions; however, most web-tools expose only a single, pre-built model with a fixed, non-configurable architecture that users cannot redesign, retrain on their own data, or compare; they are typically dedicated to one interaction type and often one species, and report prediction scores with little interpretability. These constraints force researchers across several disconnected, single-purpose tools and limit the flexibility, reproducibility, and long-term usability of existing platforms. Results: We present Deep-Interact Studio, a unified, web-based deep-learning platform that addresses these limitations by shifting interaction prediction from a model-centric to a user-driven, comparative, and interpretable paradigm. Within a single interface spanning all four interaction classes, namely protein-protein, drug-target, RNA-protein, and protein-DNA, users design their own model architectures layer by layer, configure training hyperparameters, and train them on their own data, including custom, species-specific datasets. Multiple user-built models can then be trained under identical conditions and compared side by side at both the training and inference levels, while integrated interpretability, including SHAP-based feature attribution, embedding-space visualization, and interaction hub analysis, turns predictions into auditable, mechanistically grounded results. Deep-Interact Studio is, to our knowledge, the only such platform to combine fine-grained per-layer model customization with multi-model comparison and interpretability, offering a flexible and transparent alternative to fixed, single-purpose tools.

3
F.A.D.E. (Fully Agentic Drug Engine): A Conversational AI Platform for Drug Discovery

Kantorow, J.; Mani, N.; Mohanraj, N. R.; Zong, X.

2026-06-25 biophysics 10.64898/2026.06.20.733481 medRxiv
Top 0.1%
19.3%
Show abstract

Drug discovery remains one of the costliest and most time-intensive endeavors in the pharmaceutical pipeline, with average development costs exceeding $2.3 billion per drug, timelines spanning more than a decade, and attrition rates above 90% in clinical trials. While computational methods have expanded the searchable chemical space, current pipelines remain fragmented and largely inaccessible to researchers without deep interdisciplinary expertise. Here we present F.A.D.E. (Fully Agentic Drug Engine), a multi-agent, open-source platform that converts natural language queries into potential drug candidates, substantially lowering the expertise barrier to advanced computational drug discovery. F.A.D.E. employs a three-branch hierarchical architecture that adapts to the level of available structural data for any protein target, integrating structure prediction, binding pocket detection, equivariant diffusion-based de novo ligand generation, and binding affinity estimation into a single automated pipeline. We validate F.A.D.E. on two structurally distinct targets: the epidermal growth factor receptor kinase domain (EGFR), a well-established oncology target, and cellular retinol-binding protein 1 (CRBP1), a lipid-binding protein involved in retinoid metabolism. For EGFR, our generated candidates achieved QED scores of 0.85 compared to 0.46 for the co-crystallised reference ligand, demonstrating marked improvement in predicted drug-likeness. Results across both targets confirm that F.A.D.E. can reliably generate chemically tractable, drug-like hit compounds across diverse protein classes from simple natural language input.

4
AGPI: An AI-Powered Genomic Pathogen Intelligence Platform for Integrated Classification, Visualization, and Therapeutic Targeting

Goel, A.; Mishra, P.

2026-07-11 bioinformatics 10.64898/2026.07.07.737037 medRxiv
Top 0.1%
18.9%
Show abstract

Rapid and accurate pathogen detection remains a major challenge in modern bioinformatics, as existing tools are often fragmented and require multiple specialized workflows. We present AGPI (AI-powered Genomic Pathogen Intelligence), an integrated platform that combines genomic sequence classification, biological enrichment, three-dimensional structural visualization, and AI-guided therapeutic prioritization within a single interpretable pipeline. AGPI employs a hybrid convolutional-Bidirectional Gated Recurrent Unit (BiGRU) architecture trained on DNA sequences from 40 pathogen classes spanning viruses, bacteria, fungi, and protozoan pathogens. The model achieved 99.61% validation accuracy and 94.90% accuracy on an independent held-out evaluation of 600 pathogen sequences following iterative refinement. As a proof of concept, AGPI correctly classified a Zika virus genome with 96.14% confidence, retrieved curated biological context from 245 peer-reviewed studies, and identified Ribavirin as a leading therapeutic candidate against the Zika NS5 polymerase through AI-guided molecular docking. Multi-metric ligand similarity analysis further differentiated candidate compounds according to their structural and pharmacological properties. These results demonstrate that integrated AI-driven genomic pipelines can accelerate pathogen characterization and therapeutic hypothesis generation while providing an accessible and interpretable framework for infectious disease surveillance and computational drug repurposing.

5
HalluDesign-NA: Extending HalluDesign for De Novo Nucleic Acid Design

Fang, M.; Wang, Z.; Cao, L.

2026-06-11 bioinformatics 10.64898/2026.06.10.730767 medRxiv
Top 0.1%
18.6%
Show abstract

AlphaFold3 has revolutionized the prediction of biomolecular structures and interactions, including atomic-level modeling of nucleic acids. However, the de novo design of structured and functional nucleic acids remains a significant challenge. Here, we extend our HalluDesign framework to nucleic acid design by integrating NA-MPNN for nucleic acid sequence optimization and design. This new framework, HalluDesign-NA, enables iterative sequence-structure co-optimization, facilitating the de novo design of nucleic acids. Computational benchmarking across ssDNA, ssRNA, and aptamer design tasks demonstrates consistent improvements in confidence scores (pLDDT, ipTM), supporting the feasibility of de novo nucleic acid design under various constraints, such as sequence length, symmetry, and protein structure context. We anticipate that HalluDesign-NA will accelerate the de novo design of functional nucleic acids for applications in biotechnology and medicine. The source code for HalluDesign-NA is available at https://github.com/MinchaoFang/HalluDesign_NA.

6
novelBGC: An interactive dual-score framework for biosynthetic gene cluster novelty assessment and candidate prioritisation

Shukla, G.; Merugu, B.; Sharma, G.

2026-06-18 bioinformatics 10.64898/2026.06.15.732227 medRxiv
Top 0.1%
18.2%
Show abstract

Genome mining now yields tens of thousands of putative biosynthetic gene clusters (BGCs) per project, yet, separating genuinely novel candidates from rediscoveries of known compounds remains the rate-limiting step before experimental validation. Single-axis prioritisation tools, antiSMASH similarity, BiG-FAM GCF distance, and self-resistance-enzyme (SRE) filters such as ARTS, each surface a different facet of evidence, yet their isolated use systematically over-ranks rediscovery-prone BGCs and overlooks genuinely orphan clusters. We present novelBGC, a web-hosted framework that converts these disparate outputs into two deliberately non-inverse continuous metrics per BGC, a Novelty (N) and a Reference Similarity (RS) score which together define a 2D decision plane that resolves rediscoveries, divergent family members, contig-edge artefacts, and uncharted chemistry with interactive visualisations, with all component weights user-tuneable at submission. Retrospective validation across three independent experimental datasets demonstrates the utility of the framework for candidate prioritization. Within the first 186-BGC SRE-guided cloning study, every confirmed bioactive product fell within the low-to-mid N band whereas 55 high-N (N [≥] 0.50) BGCs were never selected. Moreover, in the other two studies, it correctly prioritised the fully orphan lariocidin BGC of Paenibacillus sp. M2 and the divergent within-family indanopyrrole-A idp BGC of Streptomyces sp. CNX-425. Together, these case studies demonstrate that the joint (N, RS) space facilitates prioritization decisions that are difficult to achieve using any single criterion alone. from identical input data. novelBGC requires no command-line expertise, no local tool installation, and no manual integration of intermediate output formats, addressing a well-documented accessibility barrier for wet-laboratory researchers engaging with genome-mining workflows. novelBGC is freely available at https://project.iith.ac.in/sharmaglab/novelbgc/.

7
A Biologically Informed Heterogeneous Graph Neural Network for Multi-Task Prediction of ncRNA-Metastasis-Cancer Interactions

Midjani, F.; Shaghouzi, M.; Banadaki, A. D.; Rahimikashkooli, N.; Keshtkar, F. Z.; Malekpour, M.; Hashemi, S.; Hernandez-Barco, Y. G.; Soleymanjahi, S.

2026-08-21 systems biology 10.64898/2026.08.18.745571 medRxiv
Top 0.1%
14.3%
Show abstract

Metastasis involves context-dependent molecular interactions in which non-coding RNAs, particularly miRNAs and circRNAs, play important regulatory roles. However, existing computational approaches generally do not jointly represent cancer type, metastatic event, and cancer-specific metastatic context. We developed a context-aware multi-task heterogeneous graph neural network (GNN) for predicting ncRNA associations with cancer types and metastatic events. The framework integrates multiple biological repositories into a heterogeneous graph representing ncRNAs, cancers, metastatic event types (METs), and cancer-specific metastatic instances (CSMIs). The model performs six link-prediction tasks using a hierarchical transformer-based encoder and multi-relational TuckER decoder. Across ten independently initialized runs evaluated on the RNA-group-disjoint held-out test set, the model achieved a global AUROC of 0.8801 {+/-} 0.0118 and an F1 score of 0.8260 {+/-} 0.0071. All three ablation variants yielded lower AUROC, with the largest reduction under independent task training. Case studies in pancreatic cancer, colorectal cancer, and hepatocellular carcinoma provided disease-level, event-level, and expression-based support, respectively, for top-ranked candidate associations. The framework enables context-specific prioritization of ncRNA-cancer-metastasis associations for experimental evaluation.

8
scRepresenter: a workflow for computing, integrating and benchmarking cellular representations in single-cell transcriptomics

Pocas, G.; Umar, M.; Davis, O.; Hemberg, M.; Lamurias, A.; Lakatos, A.; Asif, M.

2026-07-20 bioinformatics 10.64898/2026.07.15.738660 medRxiv
Top 0.1%
13.1%
Show abstract

MotivationSingle-cell RNA sequencing (scRNA-seq) has become an attractive tool for studying complex diseases, in which transient cell states affecting diverse cell populations characterise disease development and progression. However, due to data sparsity and disease heterogeneity analysis is often challenging. With recent advances in machine learning, two widely used approaches have emerged for learning cellular representations: large-scale foundation models and biological knowledge-guided methods. Despite their complementary strengths, there is currently no unified workflow for systematically comparing and integrating these approaches. ResultsHere, we present scRepresenter, an open-source workflow for computing, integrating, and validating cellular embeddings derived from foundation models and biological knowledge-guided methods in the context of complex diseases. It consists of two components: a command-line workflow that computes cellular embeddings and performs downstream analyses, and an interactive Shiny application for visualizing and comparing the computed embeddings. scRepresenter supports four categories of cellular representations: (1) expression-based, (2) knowledge-guided, (3) foundation model-derived, and (4) hybrid embeddings that combine foundation model-derived representations with knowledge-guided representations. This approach takes a cell-by-gene count matrix as input and outputs an integrated object containing the computed embeddings. Then, this object can be uploaded into our interactive Shiny application to compare different embeddings. AvailabilityThe workflow is available at https://github.com/GuilhermePocas/scRepresenter ContactAL291@cam.ac.uk; MA2129@cam.ac.uk

9
pFLEX – a Python library for fast functional evaluation of genetic networks at the biological module-level

Demirtas, T.;Shaw, A.;Billmann, M.

2026-06-14 Systems Biology 10.64898/2026.06.11.731557 medRxiv
Top 0.1%
12.9%
Show abstract

Genetic networks derived from omics data are a powerful tool for systematic gene function prediction. Performance evaluation of such predictions is crucial to judge the data and computational pipeline for network construction, but unbalanced functional standards often cause hidden evaluation biases. To visualize and mitigate such biases, we previously developed the R package FLEX. Here, we present the pFLEX genetic network benchmarking tool as Python library with new and improved functionality. pFLEX improves overall runtime 4.1 to 15.8-fold. It offers additional evaluation metrics that allow for easy comparison of precision recall performance at the complex or pathway resolution between genetic networks. We demonstrate the utility of pFLEX for evaluating tissue-specific co-essentiality networks and data normalization strategies of the Cancer Dependency Map, as well as for cell line-specific Perturb-Seq-derived networks. This illustrates the requirement for biological module-resolved precision recall metrics in pFLEX for sensitive and fast evaluation of genetic networks. Availability and ImplementationpFLEX is available under the MIT license at https://github.com/billmannlab/pFLEX and the pFLEX version used in this manuscript along with benchmarking code for the analyses presented in this manuscript are archived at https://doi.org/10.5281/zenodo.20632868.

10
DNAS-Bench: Deterministic Nucleic Acid Screener Benchmarking

Wong, H. C.; Kohno, T.; Nivala, J.

2026-07-20 bioinformatics 10.64898/2026.07.06.736904 medRxiv
Top 0.1%
12.9%
Show abstract

The rapid growth of biotechnology manufacturing for synthetic DNA and proteins has raised concerns that adversaries could exploit commercial synthesis pipelines to create biological weapons. Without effective safeguards, an attacker could seek regulated genetic sequences from synthesis providers; while synthetic DNA is not itself a pathogen or toxin, access to such sequences can lower barriers to downstream misuse, motivating robust order-time screening. To mitigate this risk, Biosecurity Screening Software (BSS) systems have been developed to flag potentially malicious synthesis orders. Here, we propose one of the first deterministic benchmarks for evaluating the robustness of Biosecurity Screening Software. Our framework enables systematic testing of BSS behaviors and potential on specific nucleic-acid sequences and on targeted regions of malicious genomes. Our framework allows for insights into what is being flagged as malicious in BSSs, leading to potential discussions if specific BSS is fit for a specific manufacturing pipeline. We additionally introduce a dataset of manipulated genomes derived from the HHS and USDA Select Agents and Toxins List. When evaluated on this dataset, SeqScreen flags 42% of the sequences as malicious, while Commec flags 10.2%. Across a range of manipulation strategies, we find that simple manipulations, such as padding sequences by adding a repeated nucleotides at 1.5 times the original length, perform nearly as well as more targeted methods, such as embedding malicious sequences within benign genomic context. Padding-based methods trail embedding-based methods by only 0.75 percentage points in average detection rate. Consistent with prior reports from BSS developers and studies, we observe a sharp drop in detection rate when input sequence length falls below a critical threshold, typically between 50 and 100 base pairs (bp). Under our threat model, this implies that an adversary can bypass most existing safeguards by splitting a target genome into fragments shorter than 50 bp. Fragment-level analysis further reveals that some toxin regions evade detection entirely by SeqScreen, while other malicious genomes remain detectable even when fragmented into 30-50 base-pair segments. We open-source this benchmark to support reproducible evaluation of BSS robustness and to inform the development of next-generation biosecurity screening tools (https://github.com/HenryCWong/DNAS-Bench). For ethical concerns we only open-source the framework while the data is available upon request.

11
pylimma: a faithful, AnnData-native Python port of R limma for differential expression analysis

Mulvey, J.

2026-07-10 bioinformatics 10.64898/2026.07.06.736732 medRxiv
Top 0.1%
12.8%
Show abstract

pylimma is a faithful Python port of limma, intended to bring one of the most widely used tools for differential expression analysis to the developing Python ecosystem for transcriptomics and proteomics. We validated pylimma against the existing R implementation through 227 function-level comparisons and across six real world datasets spanning microarray, RNAseq, proteomics and single-cell transcriptomics. pylimma reproduces limmas numerical output to a median agreement of 13 significant figures and calls identical sets of differentially expressed features and gene sets. This supports its use as a drop-in replacement for the R implementation.

12
GeneAutomate: A Browser-Based, Integer-Indexed Platform for Dual-Gene-List Functional Annotation and Interactive Network Visualization

Singh, R. P.; Kumar, A.

2026-07-21 bioinformatics 10.64898/2026.07.16.738882 medRxiv
Top 0.1%
12.7%
Show abstract

Comparative interpretation of two gene lists, for example, two treatment arms, two tissues, or a discovery and a validation cohort, is a routine task in functional genomics. While several tools offer dual-list comparison (e.g., EnrichmentMap, RRHO packages), they typically require local software installation, R/Bioconductor, or manual reconciliation of separate single-list outputs. Most widely used web-based enrichment tools (DAVID, g:Profiler, Enrichr, ShinyGO, WebGestalt) are built around the analysis of a single gene list at a time, and those that support comparison often lack interactive, publication-ready visualization or depend on server-side query latency. Here we present GeneAutomate, a browser-based tool purpose-built for side-by-side comparison of two gene lists. GeneAutomate performs Over-Representation Analysis (ORA) against Gene Ontology (GO) and Reactome using an exact hypergeometric test with Benjamini-Hochberg false discovery rate correction, and Gene Set Enrichment Analysis (GSEA) when ranked (log2 fold-change) input is supplied, alongside Protein-Protein Interaction (PPI) subgraph extraction from BioGRID physical interactions. All reference data (Gene Ontology, Reactome, BioGRID, and NCBI/Ensembl identifier cross-references) are pre-compiled offline into a single integer-indexed database of approximately 32 MB for Homo sapiens, in which every gene identifier Ensembl ID, Entrez ID, official symbol, or alias is resolved to one canonical integer prior to any user query. This design removes live database round-trips from the runtime path, enabling fast, at-your-desk enrichment without installation or a server-side per-query bottleneck. The tool renders thirteen interactive, D3.js- and Cytoscape.js-based comparative visualizations, including a Rank-Rank Hypergeometric Overlap (RRHO) heatmap, a GO-slim "Radar/Spider" functional fingerprint, and chord/edge-bundled cross-talk diagrams that are, to our knowledge, not offered as an integrated set by any existing academic or commercial ORA/GSEA platform. GeneAutomate is an unfunded, individual student project developed with feedback from a professor, and is in its final stage of development. It requires no installation or login. We describe the tools architecture, statistical methods, and comparative feature set relative to established academic tools (DAVID, ShinyGO, g:Profiler, Enrichr, WebGestalt, STRING, PANTHER, GeneMANIA, Cytoscape, clusterProfiler, GSEA, Metascape) and commercial platforms (IPA, MetaCore, Pathway Studio, iPathwayGuide, Partek Pathway), and we state candidly the current versions limitations, which are planned to be the added in next version: single-species (human-only) coverage, no upstream regulator analysis, and comparison currently limited to two (occasionally three) concurrent lists. GeneAutomate is available at https://geneautomate.tech/.

13
DuplexFM: Transferable small-RNA target representations link miRNA interactions to siRNA efficacy prediction

Chen, B.; Yin, J.; Fei, J.; Yang, M.

2026-08-12 bioinformatics 10.64898/2026.08.06.743413 medRxiv
Top 0.1%
12.6%
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWMicroRNAs (miRNAs) and small interfering RNAs (siRNAs) share Argonaute-mediated guide-target recognition, yet quantitative siRNA efficacy measurements are substantially scarcer and more costly to generate than miRNA-target interaction data. We therefore asked whether miRNA interaction data could provide transferable supervision for siRNA efficacy prediction. Here we present DuplexFM, a biologically grounded framework that uses sample-specific gates to integrate five evidence sources: pairing and sequence-context priors, duplex energetics, experimentally supervised mRNA accessibility, target-to-guide cross-attention, and contextual token-pair compatibility. The accessibility expert, trained on nucleotide-resolution icSHAPE measurements, achieved a held-out nucleotide-level Pearson correlation of 0.627 and evaluated accessibility at seed match and energy-supported candidate sites. On miRBench v7, three independently trained DuplexFM models achieved a macro APS of 0.873{+/-}0.002, soft-voting increased this to 0.876 and yielded the highest APS on all four test sets. We then froze the miRNA-trained representation and trained only a lightweight residual head with 24 siRNA-specific descriptors. Transfer improved Pearson and Spearman correlations, AUPRC, and F1 over the descriptor-only baseline in all six evaluation settings. The ensemble achieved the highest Pearson and Spearman correlations in four settings, whereas OligoFormer remained stronger on Huesken and Takayuki. These findings show that experimentally grounded accessibility and miRNA-derived interaction representations provide complementary, transferable information, supporting a parameter-efficient route towards unified modeling of Argonaute-guided RNA regulation. Code and data are available at https://github.com/cbaiming/DuplexFM.

14
An Integrated Knowledge Graph and Network Medicine Pipeline for Drug Repurposing: Benchmarking Across Human Diseases and Application to Amyotrophic Lateral Sclerosis

Jiang, A.; Hu, J.; Abdulle, Y.; Pain, O.; Iacoangeli, A.

2026-07-08 bioinformatics 10.64898/2026.07.03.736387 medRxiv
Top 0.1%
12.4%
Show abstract

Drug repurposing offers a practical strategy to identify new therapeutic uses for approved drugs, potentially reducing the time and cost associated with conventional drug development. We present a novel three-stage drug repurposing pipeline that integrates knowledge graph-based gene prediction, network-based drug-disease association analysis, and systematic classification of candidate drugs by therapeutic class. The pipeline integrates DGLinker to predict novel disease-associated genes, SAveRUNNER to identify drug repurposing candidates, and ATC Category Enrichment Analysis (ATCEA) to prioritise candidates by pharmacological class. We benchmarked the pipeline across twelve diseases using DrugBank and MEDI2-HPS as validation resources. Utilising DGLinker-expanded disease-gene sets as input increased the number of predicted repurposed drugs, while overall discriminative performance remained stable across diseases (AUROC 0.71-0.77). Application of ATCEA consistently improved precision, F1-score, and specificity, while reducing recall, reflecting a conservative prioritisation strategy that contracts the candidate space while retaining pharmacologically coherent drug-disease candidates. We further applied the pipeline to amyotrophic lateral sclerosis (ALS), a neurodegenerative disease with limited therapeutic options, and performed a deeper literature-based validation of the results. Incorporation of DGLinker-predicted genes substantially increased the number of significant candidate drugs and uncovered enriched ATC categories not identified using known ALS genes alone, including antidepressants and antipsychotics. Moreover, several drugs with supporting evidence available in the literature were identified only when DGLinker-predicted genes were used. Overall, 77 candidate drugs were prioritised within significantly enriched ATC categories, several of which are supported by previously published studies. To provide exploratory real-world support for these findings, we further evaluated candidate drugs in a longitudinal electronic health record (EHR) dataset of 2361 patients with ALS from King's College Hospital. Although the number of evaluable drugs was limited due to sample size, the EHR analysis provided additional clinically relevant context for selected prioritised drugs and pharmacological classes. Our pipeline demonstrates potential to accelerate drug repurposing by integrating complementary computational approaches to each step of the process, providing an end-to-end framework that showed robust performance across benchmarking experiments and use cases.

15
EpiESM-GA: Resource-Efficient Protein Foundation Model Features for Equitable B-Cell Epitope Prediction

Gautam, P.; Mitra, P.

2026-06-26 bioinformatics 10.64898/2026.06.22.733745 medRxiv
Top 0.1%
12.0%
Show abstract

Prediction of B-cell epitopes can assist in reducing costly wet-lab screening in vaccine design, diagnostics, and antibody discovery. However, current predictors often suffer from noisy labels, weak generalization, and structure-dependent workflows. Here we present EO_SCPLOWPIC_SCPLOWESM-GA, an efficient sequenceonly pipeline for linear B-cell epitope prediction. Positive and negative peptide examples are collected from IEDB, which provides experimentally tested epitopes and distinguishes positive and negative epitope records based on assay evidence(Vita et al., 2019). Each peptide is encoded with a frozen ESM-2 protein language model: a bidirectional transformer producing amino acid embeddings for downstream structure and function tasks (Lin et al., 2023). Mean-pooled embeddings are further compressed into a compact 420-feature representation with a genetic algorithm and classified with lightweight Random Forest, XGBoost, or MLP heads. This avoids foundation-model fine-tuning, reduces the number of trainable parameters, improves interpretability, and enables low-resource deployment. On an IEDB-derived benchmark, EO_SCPLOWPIC_SCPLOWESM-GA attains 0.880{+/-} 0.004 AUC-ROC, 0.852{+/-} 0.005 PR-AUC, 82.0 {+/-} 0.6% accuracy, 0.79 {+/-} 0.01 F1, and 0.74{+/-} 0.01 MCC, outperforming dense ESM-2 features and baselines LBCE-XGB, EpitopeVec, and BepiPred-2.0 (mean{+/-} std over five independent random seeds). The framework shows how frozen protein foundation models can enable pandemic preparedness, peptide vaccine prioritization, diagnostic antigen screening, and equitable computational immunology.

16
Benchmarking Graph Neural Networks for Multi-Omics Cancer Subtyping using Methylation and Gene Expression Profiles

Schirmacher, J.; Maurer, M. C.; Metsch, J. M.; Ploesch, S.; Chereda, H.; Blumenthal, D. B.; Hauschild, A.-C.

2026-08-25 bioinformatics 10.64898/2026.08.21.745839 medRxiv
Top 0.1%
12.0%
Show abstract

Motivation: Graph Neural Networks (GNNs) have gained increasing interest in the biomedical domain, as the integration of prior knowledge and deep neural networks has the potential to enhance insights into molecular processes and disease mechanisms. However, a comprehensive and systematic assessment of model architectures, data modalities, graph structures, and their performance for graph signal classification in the biomedical domain is yet to be performed. In order to close this gap, we conducted a benchmarking study on multiple GNNs on a Protein-Protein Interaction (PPI) network for Kidney Renal Clear Cell Carcinoma and Breast cancer subtype prediction, performing an in-depth investigation of architectures, incorporating skip connections and various data modalities. Results: While none of the GNNs outperforms the structure-agnostic Multi-Layer Perceptron baseline, all of them can handle bimodal data (gene methylation and expression) and offer the ability to gain explainability based on PPIs. We offer practical guidelines for applying GNNs to graph signal processing tasks specifically for cancer classification. Depending on the underlying dataset and PPI structure employed, models on different data modalities outperform others. Overall, we suggest using ChebNet, which tends to outperform the Graph Convolutional Network and the Graph Attention Network in cancer subtype prediction. We recommend using GNN architectures that employ a simple flattening readout layer, as they provide better classification performance and faster training time than those with global average pooling. Additionally, we tested residual connections, but they had only an insignificant impact on classification performance.

17
A Structural Antibody Benchmark of AlphaFold3 reveals Hallucinated Epitopes and a Bias for Orderness

Solanki, A.; Maurya, N. S.; Ramlakhan, M.; Li, R.; Chen, W.; Wu, Z.; Zheng, W. J.

2026-07-31 bioinformatics 10.64898/2026.07.30.741792 medRxiv
Top 0.2%
11.8%
Show abstract

AlphaFold3 has shown promise as a tool for predicting antibody-antigen binding, yet its performance across large datasets has not been fully characterized. In this study, 3401 experimentally validated antibody-antigen complexes were sourced from the Structural Antibody Database and screened alongside 23798 negative controls to benchmark AlphaFold3s binding prediction capabilities. Confidence metrics including Predicted Aligned Error and Interface Predicted Template Modeling score were used to achieving a maximum recall of 53% at 100 inference seeds. Several factors were found to influence prediction accuracy: a notable bias was observed toward antibodies derived from X-ray crystallography structures versus those from electron microscopy, and positive prediction rates were found to decrease with increasing target protein size and surface area. In contrast, neither the amino acid composition or lengths of the complementarity determining regions, nor training data leakage were found to introduce significant bias. An innate false positive rate of approximately 3% was identified, with AF3 shown to hallucinate plausible binding interfaces across the surface of decoy targets while avoiding disordered regions. Epitope mapping using DockQ, epitope shift, and antibody displacement revealed that approximately 34% of false negatives retained the correct epitope location despite poor structural alignment, suggesting that conformation refinement tools could recover additional true binding predictions. These findings provide a comprehensive characterization of AlphaFold3s strengths and limitations for antibody screening in computational drug discovery. Key MessagesO_LIAlphaFold3 has a recall of 50% and an innate false positive prediction rate of 3%. C_LIO_LIFalse negative predictions can still feature the correct epitope despite poor RMSD. C_LIO_LIFactors such as disorder and target size impact accuracy. C_LI

18
CyChat: a conversational Cytoscape app for no-code, reproducible network analysis

Liebold, J.; Stahl, M.; Schulze, J.-O.; Razavi, M. M.; Bader, G. B.; Kurtz, S.; Baumbach, J.

2026-09-01 bioinformatics 10.64898/2026.08.28.747833 medRxiv
Top 0.2%
11.8%
Show abstract

Network-based analyses of molecular interactions are useful for interpreting high-throughput omics data and identifying therapeutic targets. Cytoscape is the standard platform for these tasks, but users face a trade-off between accessible graphical workflows that are difficult to document and reproducible automation in Python or R that requires programming expertise. General-purpose coding assistants can generate Cytoscape Automation scripts, but remain external to Cytoscape. We present CyChat, a Cytoscape Desktop app that integrates a chat interface and a large language model (LLM) agent into the application. CyChat translates natural language into executable Cytoscape Automation workflows, runs generated Python code, and exports chat sessions with executed code as standalone Jupyter notebooks. To reduce setup barriers, CyChat includes an embedded Python runtime and supports both cloud-based and locally hosted LLMs. CyChat was evaluated across ten Cytoscape workflows using seven LLM providers, each represented by one LLM. The strongest configuration achieves a pass rate above 99%. In a qualitative evaluation based on a published network visualization, CyChat completes the task in 1.5-5 minutes, compared with 15-20 minutes for manual GUI workflows by computational biologists. CyChat is available through the Cytoscape App Store at https://apps.cytoscape.org/apps/cychat.

19
scINTILLA: Single-Cell Integrated Inference, Labelling, and Landscape Analysis for Cell-Type Annotation Quality Assessment

Kanannejad, S.; Bongiorni, N.; Nordera, E.; Redaelli, S.; Rusconi, I.; Zanin, R.; Giustacchini, A.; Chatterjee, S.

2026-07-28 bioinformatics 10.64898/2026.07.27.740477 medRxiv
Top 0.2%
11.6%
Show abstract

Single-cell RNA sequencing has enabled the construction of comprehensive cell atlases, yet the quality and coherence of the cell-type annotations within these atlases remain largely unexamined. When a label is applied to a transcriptionally heterogeneous population, the downstream analyses that depend on it, and automated label transfer in particular, become unreliable. We present scINTILLA (Single-Cell Integrated Inference, Labelling, and Landscape Analysis), a computational framework that combines supervised and unsupervised machine learning to score the learnability and internal consistency of cell-type labels in single-cell datasets. The unsupervised arm benchmarks a broad panel of clustering algorithms and derives a neighbourhood confusion score for every cell, whilst the supervised arm trains up to twelve classifiers and extracts prediction agreement, entropy, and confidence. These signals are normalised and aggregated into a single composite score per cell type, where a low score flags label ambiguity or concealed heterogeneity. As a by-product, scIN-TILLA also reports which clustering and classification algorithms perform best on a given dataset, offering practical guidance for downstream label transfer. We applied it to five Human Cell Atlas datasets spanning the adult brain, lung, eye, and two organoid atlases, and recovered clear differences in the learnability and internal consistency of annotations across atlases that were not driven by the number of annotated cell types. Focused re-analysis of lowscoring populations in the lung and endoderm-organoid atlases resolved biologically coherent sub-populations, in some cases with context-specific enrichment, much of it recovered from cells that had been assigned broad or catch-all labels. scINTILLA is advisory rather than prescriptive, guiding principled, data-driven re-annotation at atlas scale.

20
Tangerine: A Python framework for dynamic gene regulation analysis from transcriptomic time series

Narendra, T.; Schweikert, G.

2026-07-22 bioinformatics 10.64898/2026.07.17.739167 medRxiv
Top 0.2%
11.6%
Show abstract

MotivationTime-series single-cell transcriptomics enables the study of dynamic gene regulation. However, standard computational tools frequently aggregate temporal data into static, dense topologies, obscuring the precise regulatory rewiring that drives developmental transitions. Further, navigating the inherent noise of statistical inference without losing biological interpretability remains an important bottleneck. ResultsWe present Tangerine, a Python framework for the dynamic reconstruction and interactive exploration of time-varying gene regulatory networks. Tangerine integrates time-constrained metacell aggregation with regularized linear modelling and non-parametric correlation to infer dynamic topologies. To solve the interpretability gap, it features a browser-based visual analytics engine. Tangerine empowers researchers to track macroscopic gene module evolution, interactively filter effect sizes, and link topological rewiring directly to raw transcriptomic evidence. Availability and implementationTangerine is implemented in Python and Plotly Dash. The code is available on Github at https://github.com/ntanmayee/tangerine.